Medical Physics
○ Wiley
Preprints posted in the last 30 days, ranked by how well they match Medical Physics's content profile, based on 14 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Oyarzun Silva, R.; Hernandez Hernandez, P.
Show abstract
Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.
Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.
Show abstract
Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.
dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.
Show abstract
Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.
Sandvold, O. F.; Proksa, R.; Perkins, A. E.; Daerr, H.; Koehler, T.; Jacob, T.; Brown, K. M.; Roessl, E.; Noël, P. B.
Show abstract
Spectral computed tomography (CT) is a burgeoning quantitative imaging technique with applications in oncologic diagnostics, prognostic prediction, tissue perfusion studies, and treatment follow-up. While normalized iodine concentration values have been correlated with microenvironmental biophysical changes, obtaining accurate iodine concentrations, particularly at low concentrations remains difficult due to varying spectral CT instrumentation performance. Hybrid spectral CT systems, combining multiple spectral CT instrumentation techniques, address these quantitation insufficiencies by increasing spectral separation but have not been evaluated on a clinically analogous platform. We validate a hybrid spectral CT system, comprised of clinical-grade components, acquiring four distinct effective spectra and applying efficient noise-reducing weighting schemes to compare iodine noise and bias against conventional kVp-Switching (kVp-S). Two tube current levels (50, 350 mA) and three duty cycle ratios (33/67, 50/50, 75/25) were implemented to elucidate radiation dose exposure and kVp-S parameterization impact. A standard quality assurance (QA) and patient-derived, abdominal IodinePrint phantom were scanned on the system. The average absolute bias in iodine density images of the QA phantom was comparable across acquisition techniques, below 0.5 mg/mL, while quantitative noise improved by 22% using noise-optimized weighting schemes. In the IodinePrint phantom aorta and pancreas structures, the noise-optimized weighting scheme increased signal-to-noise ratio (SNR) by 1.3x compared to kVp-S alone. These results highlight the increased precision of hybrid, multi-channel spectral CT systems and motivate CT designs that enable robust CT biomarker development.
Do, H. P.; Bekku, M.; Berkeley, D.; Golden, M.; Kitane, S.; Uike, M.; Shinoda, K.; Takayanagi, R.; Takai, H.; Kawai, T.; Seballos, K.; Conley, R.; Sorfleet, K.; Devries, D.; Tymkiw, B.; AlGhuraibawi, W.; Caruthers, S. D.; Kadbi, M.; Provencher, M.; Tashman, S.; Ho, C. P.
Show abstract
Purpose: To determine the feasibility of a 2-minute multi-echo UTE (mecho-UTE) for CT-like bone-weighted contrast and T2* quantification of tissues with short T2/T2*. Methods: Mecho-UTE data acquired from four patients and five healthy subjects were used to assess image quality of the CT-like contrast. All data were reconstructed using conventional gridding (GRID+CONV) and compared with those reconstructed using conjugate gradient SENSE combined with deep learning-based denoising (CG+DLR). Image resolution and sharpness of the CT-like images were assessed using the full width at half maximum (FWHM) and relative edge sharpness (RESH), respectively. Calimetrix UTE-T2* phantom was used to assess the accuracy of T2* quantification of the mecho-UTE sequence. Results: Two-minute mecho-UTE with CG+DLR has similar accuracy (0.37 {+/-} 0.27 vs. 0.67 {+/-} 0.54 ms, p=0.20) and better precision (0.28 {+/-} 0.16 vs. 1.23 {+/-} 0.29 ms, p<0.001) compared to the 5-minute mecho-UTE with GRID+CONV. The 2-minute mecho-UTE with CG+DLR has higher resolution and sharpness compared to the 5-minute scan with GRID+CONV. Conclusion: It is feasible to achieve simultaneous CT-like contrast and T2* quantification of short-T2 tissues in two minutes. When appropriately used, it may simplify logistics, reduce costs, and eliminate radiation exposure risks.
Zareian, B.; Fontaine, K.; Bini, J.
Show abstract
Background. Roughly, half of new type 1 diabetes (T1D) diagnoses occur in individuals under 18 years old and represent a more aggressive destruction of beta cell mass (BCM). [11C]-(+)-PHNO positron emission tomography (PET) imaging is used to assess BCM, but current pancreas PET imaging protocols are limited to adults. Previously published full count data from six healthy controls and five T1Ds (6M/5F; 22 to 53 years old) were used for retrospective analysis. Dynamic [11C]-(+)-PHNO PET/CT scans were acquired and reconstructed using full-count list-mode data. For the current comparison to full count data, 50%, 25% and 10% down-sampled count data were re-reconstructed. Pancreas and spleen (reference region) time-activity-curves (TACs) were assessed, and volume of distribution (VT, mL/cm3) was estimated using the reversible 1-tissue compartment model (1TC) with tmax of 30 min for all count levels. Pancreas and Spleen VT estimates (1TC; tmax= 30 min) were used to calculate non-displaceable binding potential (BPND) and were then correlated to semi-quantitative methods of standardized uptake value ratio (SUVR-1) (20-30 min; ref: spleen) to examine simplified methods using simulated low dose protocols. Finally, we performed dosimetry in adult, adolescent and pediatric phantoms to assess radiation dose for simulated low-dose protocols. Results. Qualitatively, increasing noise can be visualized at successive reduced-count levels images, compared to full-count images. Despite progressively increasing noise in reduced-count images, TACs at each reduced-count level remained similar to full-count TACs in both HC and individuals with T1D. Quantitatively, 1TC VT estimates were similar for all reduced count levels and range of tmax values, compared to full-count (all R2[≥]0.99). Pancreas SUVR-1 (20-30 min) and pancreas BPND (tmax = 30; ref: spleen) were highly correlated for all count levels (all R2[≥]0.80). All age groups were under both the yearly occupational and research scan radiation dose limits when examining mean effective dose equivalent with reduced (1/10th) injected dose protocols. Conclusion. Low-count reconstructed data and simplified reference region approaches provide accurate quantification compared to full-count reconstructions. These results provide evidence that it is possible to perform accurate quantification using simulated low dose protocols to quantify BCM for use in individuals with T1D under 18 years old.
Suzuki, M.
Show abstract
Background. Extracellular volume fraction (ECV) derived from contrast-enhanced CT is a validated marker of hepatic fibrosis and has been reported to differ between hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma. In published work it is obtained from a small number of hand-placed two-dimensional regions of interest, and the software that computes it is either tied to one manufacturer's workstation or based on spectral or dual-energy acquisition. We are not aware of an accessible tool that produces voxelwise liver ECV maps from conventional single-energy multiphase CT. Methods. We developed CT ECV Mapper, a scripted 3D Slicer extension with a three-layer architecture whose numerical core imports neither slicer nor vtk and is unit-tested outside 3D Slicer. The interactive application provides two-stage registration that the operator inspects and accepts before any ECV is computed, operator-placed three-dimensional regions of interest, user-adjustable calculation parameters, a voxelwise ECV color map and ROI statistics; the same logic layer can be driven unattended across a cohort. The tool was applied to the 164 patients of the public WAW-TACE multiphase HCC/TACE dataset that have both unenhanced and delayed-phase series. Results. 156 of 164 cases (95.1%) completed unattended. Whole-liver ECV had a median of 36.2% (interquartile range 31.9-41.5), consistent with published CT-ECV values for fibrotic and cirrhotic liver. Registering the arterial and portal phases on demand extended tumor ECV from the 38 lesions a conventional two-phase pipeline can reach to 248 lesions in 156 patients. Every failure was attributable to an identifiable mechanism: craniocaudal field-of-view mismatch between phases in six cases, aortic calcification within the blood-pool region in one, and in one case a labeling error in the source dataset, in which the series declared as unenhanced proved to be a second reconstruction of the portal venous phase; this was detected by the blood-pool validity check rather than by visual review. Conclusions. Voxelwise CT ECV mapping of the liver and of hepatic tumors is feasible from conventional multiphase CT on an open platform, both interactively and as an unattended batch, with quality-control instrumentation that fails explicitly and diagnosably. This is a technical development and feasibility report; the application has not been evaluated against a reference standard and no claim of clinical validity is made.
Lan, W.; Weigel, S.; Calderon, E.; Fougere, C. l.; Schmidt, F. P.
Show abstract
Purpose: Respiratory motion remains a major source of quantitative bias in PET and becomes increasingly relevant for high-sensitivity long axial field-of-view (LAFOV) PET/CT. Although numerous respiratory motion correction (MoCo) methods have been proposed, their quantitative accuracy cannot be established clinically because a patient-specific motion-free reference is fundamentally unavailable in vivo. This study combined clinical PET imaging with a digital twin, a realistic representation of both the PET/CT system and the patient, to objectively validate respiratory MoCo against a corresponding motion-free reference. Methods: Twenty patients (10 [18F]FDG with predominantly pulmonary lesions and 10 [18F]SiFAlin-TATE with predominantly hepatic lesions; total 135 lesions) were analyzed. The digital twin combined a validated LAFOV PET/CT simulation model with an anatomically realistic phantom containing 14 lung and liver lesions, two patient-derived respiratory patterns, and respiratory motion amplitudes of 2 and 3 cm, generating patient-like datasets with corresponding motion-free references. Data-driven and image-based MoCo were evaluated using lesion morphology, SUVmean, SUVmax, and metabolic tumor volume (MTV). Results: In patients, data-driven MoCo produced larger SUVmean increases than image-based MoCo for liver (48.1{+/-}18.9% vs. 17.0 {+/-} 12.0%; p<0.01), lower-lung (32.5{+/-}21.2% vs. 16.3{+/-}15.6%, p=0.06), and upper-lung lesions (28.4{+/-}32.0% vs. 10.4 {+/-} 17.2%; p<0.01), with similar findings for SUVmax and larger MTV reductions. Simulation revealed marked motion-induced SUVmean underestimation before correction, particularly in liver (-31.2{+/-}6.8%) and lower lung (-15.5{+/-}13.9%). Relative to the motion-free reference, data-driven MoCo most accurately recovered hepatic uptake (4.3{+/-}11.7% vs. -10.0 {+/-} 9.2%; p=0.01) but overestimated pulmonary uptake (lower lung: 19.8{+/-}16.3% vs. -1.6 {+/-} 10.2%; p=0.02). SUVmax showed the same regional behavior, whereas image-based MoCo yielded MTV estimates closer to the reference. Quantitative recovery was largely independent of respiratory pattern, while larger motion amplitudes mainly affected image-based MoCo. Conclusion: Combining clinical PET with a realistic digital twin and corresponding motion-free ground truth enabled objective validation of respiratory MoCo beyond conventional clinical evaluation. Larger correction-induced quantitative changes should not be equated with greater quantitative accuracy. Instead, MoCo performance was region- and metric-dependent, highlighting the value of ground-truth-based validation for developing and benchmarking respiratory motion correction and quantitative PET on LAFOV PET/CT systems.
BAI, T.-C.; YEH, S.-C.
Show abstract
Foundation models for chest X-ray interpretation make it possible to adapt specialised visual representations with relatively small trainable modules. We report a retrospective study of Low-Rank Adaptation (LoRA) of Rad-DINO Vision Transformer Base with 14x14 patches (ViT-B/14) for 14-class multi-label classification on the National Institutes of Health (NIH) ChestX-ray14 dataset. The official test labels were accessed during earlier model development and configuration comparisons; consequently, every official-test result in this manuscript is explicitly descriptive and non-confirmatory. We used a patient-disjoint 90/10 split of the official trainval pool (77,988 training and 8,536 validation images) and retained the released 25,596-image test partition. The historically selected all-linear LoRA configuration with safe augmentation and g=37 produced a descriptive test macro AUROC of 0.8462 versus the frozen baseline of 0.8295. Comparisons of target modules, patch-token grids, and a Rad-DINO-specific local query head are reported as retrospective comparisons rather than unbiased model-selection evidence. A confident-learning diagnostic flagged 17,653 of 86,524 trainval images (20.4%); this is a model-based flag rate, not a ground-truth label-error rate. A separate counterfactual relabeling sensitivity analysis, which uses the same model to identify and rescore disagreements, changed the descriptive AUROC to approximately 0.9445 after 6,509 policy-defined flips. This value is not achieved model performance and is not a radiologist-audited label-quality ceiling. We provide a validation-only threshold and artifact protocol for future locked evaluation, but a genuinely untouched holdout and new locked selection are required for a confirmatory headline. The existing Zenodo record contains the 25 publication figures only.
Hong, V.; Bulent, A.; Haouchine, N.; Pieper, S.; Wells, S.; Keko, M.; Kozono, D.; Doyle, P. F.; Balboni, T.; Spektor, A.; Huynh, M. A.; Hackney, D. B.; Alkalay, R. N.
Show abstract
Purpose: Clinical assessment of vertebral lesion quality (osteolytic, osteoblastic, mixed) remains subjective, with limited interobserver reliability. This study evaluated a novel application of 3D convolutional neural networks (3D-CNNs) for classifying lesion quality from CT volumes in metastatic cancer patients. Materials and Methods: This retrospective study used CT data from 151 cancer patients planned for radiotherapy for metastatic spine disease (September 2020-July 2024). Leveraging vertebra-level expert annotations, we introduced an unconventional U-Net-based strategy converting coarse voxel-wise predictions into vertebra-level lesion classifications. The final dataset comprised 2,125 vertebrae across four classes (no lesion, osteolytic, osteoblastic, mixed), split into a 3-fold cross-validation set and an independent holdout test set. Model performance was benchmarked against a DenseNet121 baseline and a musculoskeletal radiologist, with Cohen's kappa assessing inter-rater agreement. Results: The 3D model achieved an ensemble accuracy of 84.7%, outperforming DenseNet121 (72.1%), with substantial gains in F1 score, precision, and balanced accuracy. It showed high concordance with the radiologist (Cohen's kappa = 0.76) and comparable sensitivity and specificity across all lesion subtypes. We found both models and the radiologist to struggle with osteolytic lesions, reflecting the difficulty of distinguishing this class from age-related changes in vertebral bone density and architecture caused by benign bone lesions, age-related systemic skeletal disorders and cancer treatments. Conclusions: 3D-CNNs trained with vertebra-level labels can accurately and reliably classify vertebral metastatic lesion quality from CT scans, offering a scalable path toward automated characterization of metastatic spine disease to support clinical decision-making and large-scale radiomics research.
Calado, A.; de Almeida, J. G.; Verde, A. S. C.; Tsiknakis, M.; Marias, K.; Regge, D.; Papanikolaou, N.; ProCAncer-I Consortium,
Show abstract
Purpose: To prospectively validate a semi-supervised learning framework with a lesion-only teacher model (RG-SSL-LOC) for scalable clinically significant prostate cancer detection on biparametric MRI (bpMRI) and assess its added value in multimodal models. Materials and Methods: A multicenter dataset of 13,706 bpMRI examinations (13,630 patients, 27 centers) was used for model development/validation. Three segmentation models (fully supervised learning [FSL], a state-of-the-art report-guided semi-supervised approach [RG-SSL], and the proposed RG-SSL-LOC) were evaluated at lesion- and case-level on external retrospective, external prospective, and internal prospective cohorts. Predictions from the best-performing model were combined with clinico-radiologic variables in a multimodal approach. All case-level results were compared with PI-RADS. Results: At lesion level, RG-SSL-LOC achieved higher median Dice than FSL and RG-SSL (0.49 vs 0.41 and 0.40; both p<.001). At case level, RG-SSL-LOC achieved area-under-the-curve (AUC) values of 0.83, 0.82, and 0.87 in the external retrospective, external prospective, and internal prospective cohorts, respectively. Compared with FSL, AUCs were 0.84 (p=.237), 0.80 (p=.020), and 0.84 (p<.001); compared with RG-SSL, AUCs were 0.83 (p=.929), 0.82 (p=.652), and 0.86 (p=.007); compared with PI-RADS, AUCs were 0.78 (p=.055), 0.83 (p=.652) and 0.86 (p=.480). Combined with clinico-radiological variables, RG-SSL-LOC significantly improved AUC versus clinico-radiological variables alone in the external retrospective (0.85 vs 0.80, p=.002), external prospective (0.87 vs 0.84, p=.008), and internal prospective (0.91 vs 0.88, p<.001) cohorts; in the latter, it reduced unnecessary biopsies by 15.19%. Conclusion: RG-SSL-LOC achieves better segmentation quality than other methods, demonstrates robust prospective multicenter performance and improves multimodal detection.
Yi, J.; Patel, K. K.; Miller, R. J. H.; Marcinkiewicz, A. M.; Kamagate, A.; Shanbhag, A.; Hijazi, W.; Lemley, M.; Zhou, J.; Liang, J. X.; Ramirez, G.; Mostafavi, S.; Urs, M.; Spielvogel, C. P.; Slipczuk, L.; Travin, M.; Alexanderson, E.; Caraval-Juarez, I.; Packard, R. R.; Al-Mallah, M.; Ruddy, T. D.; Einstein, A. J.; Feher, A.; Miller, E. J.; Acampa, W.; Knight, S.; Le, V. T.; Mason, S.; Calsavara, V. F.; Chareonthaitawee, P.; Wopperer, S.; Kwan, A. C.; Wang, L.; Li, D.; Fishman, E. K.; Lopez-Ramirez, F.; Berman, D. S.; Kwiecinski, J.; Dey, D.; Di Carli, M. F.; Slomka, P.
Show abstract
Background: Body composition is recognized as a major determinant of health outcomes, but its multidimensional nature makes clinical adoption challenging. We sought to develop and validate a body composition index (BCI) for all-cause mortality risk assessment, integrating variables of six body composition tissues. Methods: We analyzed 28509 consecutive patients undergoing myocardial perfusion imaging with routine low-dose chest CT attenuation correction (CTAC) scans acquired during myocardial perfusion imaging (MPI) at 12 centers across four countries. An artificial intelligence-based BCI was developed in a cohort of 15037 patients CTACs by integrating the CT-derived metrics of bone, skeletal muscle, and four adipose tissue compartments, coronary artery calcium score, and basic demographic variables (age, sex, BMI). The performance of BCI for mortality prediction was validated in an internal cohort of 6444 patients and an external cohort of 7028 patients by prognosis, calibration, net benefit, and explainability. Model-based simulation of tissue metrics modification was performed to evaluate estimated mortality risk reduction. Findings: During a median of 3.5 (IQR [1.9, 5.1]) years, 4697 (16%) patients died. In the external testing cohort, the BCI demonstrated excellent discrimination for mortality (area under receiver operating characteristic curve 0.78 (95% CI [0.76, 0.79]) and Harrell concordance index 0.75 [0.73, 0.76]), calibration, and net benefit overall and across pre-specified subgroups stratified by patient characteristics and imaging protocols. Visceral adipose tissue attenuation was the most influential body composition measure, followed by skeletal muscle volume. Simulated improvement in body composition was associated with significant mortality risk reduction. Interpretation: An index combining six body composition measures obtained opportunistically from routine chest CT provides robust mortality risk stratification. By converting complex body composition information into a single interpretable score, the BCI can facilitate clinical implementation of opportunistic CT biomarkers and guide individualized preventive strategies.
BAI, T.-C.; YEH, S.-C.
Show abstract
CXR report generation may require a vision-language model (VLM) to produce both textual findings and spatial bounding boxes. Generative 4B-7B VLMs can emit non-empty outputs on normal images and empty outputs on abnormal images, motivating explicit structural routing. To evaluate whether a hard inference-time gate before a probabilistic VLM changes output-presence performance and to identify the mechanisms underlying paired STRUCT outcomes. We evaluated CXRxVLM v2, combining a frozen microsoft/rad-dino ViT-B/14 encoder with a 768[->]1 logistic probe (threshold 0.0557) and google/medgemma-4b-it with the pamessina/medgemma-4b-it-cure LoRA adapter. A seed=42 stratified cohort of 500 VinDr-CXR train-pool images (250 NORMAL, 250 ABNORMAL) was compared with Lingshu-7B A_baseline and D_fewshot configurations. Exact paired McNemar tests and stratum-level output-presence analyses were prespecified for the primary configurations; MedGemma 1.5 SigLIP was exploratory. CURE achieved STRUCT = 78.0% (390/500; Wilson 95% CI 74.2-81.4), versus 73.8% for Lingshu A_baseline and 74.2% for D_fewshot. Pairwise p-values were 0.0778, 0.1042, and 0.8642. The paired decomposition showed CURE ABNORMAL non-empty-output advantage of +13.6 percentage points versus Lingshu A (p = 0.0012; +14.0 points versus D, p = 0.0007), while Lingshu had higher NORMAL empty-output rates (+5.2 to +6.4 points; p = 0.0106 and p = 0.0004). The full pipeline used 8.87 GB VRAM and 4.92 s/image mean latency; 53% of records used a 25.7 ms warm gate-negative path after model loading. Equivalent overall STRUCT scores concealed two mechanistically different output regimes: CURE favored ABNORMAL non-empty outputs, whereas Lingshu favored NORMAL empty outputs. This paired decomposition, rather than the aggregate score alone, characterizes how hard-gated and probabilistic systems route output presence.
Bagchi, R.; Yee, N. J.; Kwon, J. Y.; Taseh, A.; Ashkani-Esfahani, S.
Show abstract
Purpose To evaluate whether domain-adaptive self-supervised pretraining on musculoskeletal radiographs improves fracture classification and attribution faithfulness relative to ImageNet-pretrained baselines. Materials and Methods This study (June 2025 to May 2026) used previously acquired radiographs to compare three ResNet-50 initializations: supervised ImageNet pretraining (control), self-supervised ImageNet pretraining (DINO), and DINO with additional domain-adapted pretraining on 44,029 musculoskeletal radiographs (DINO-Ortho). All models underwent supervised fine-tuning in three experiments: in-distribution (MURA and FracAtlas datasets), out-of-distribution (an external dataset of 5,365 calcaneal radiographs from 1,775 patients), and initial weights (calcaneal radiographs only). Metrics included sensitivity, specificity, test accuracy, area under the receiver operating characteristic curve (AUROC), and Cohen's kappa; attribution faithfulness was quantified using Remove and Debias scores from Grad-CAM saliency maps. Comparisons used DeLong and Friedman tests. Results Classification performance did not differ significantly between DINO-Ortho and either baseline in any experiment (DINO-Ortho AUROC, 0.89 in-distribution and 0.95 with initial weights). All three models discriminated poorly out-of-distribution (control, 0.59; DINO, 0.57; DINO-Ortho, 0.58). DINO-Ortho showed significantly higher attribution faithfulness than both baselines in all three experiments, including out-of-distribution (25.39 vs -10.41 and 2.14; P < .001) and initial weights (20.88 vs 11.51 and 1.27; P < .001). Qualitative rankings favored DINO-Ortho but did not differ significantly. Conclusion Domain-adapted self-supervised pretraining on musculoskeletal radiographs improved attribution faithfulness while maintaining classification performance comparable to ImageNet-pretrained baselines; no model generalized adequately to external radiographs without task-specific fine-tuning.
Segi, N.; Okada, Y.; Takeichi, Y.; Ito, S.; Ouchida, J.; Nagatani, Y.; Kagami, Y.; Tachi, H.; Ohshima, K.; Ogura, K.; Imagama, S.; Nakashima, H.
Show abstract
Study design Retrospective cohort study. Objectives To correlate Hounsfield unit (HU) values, using elliptical regions of interest (ROI), that can be easily defined in routine clinical practice with magnetic resonance imaging (MRI) T2-hyperintense area fraction, as a surrogate for paraspinal muscle fat infiltration and to establish specific HU screening thresholds that may be applied with standard picture archiving and communication system (PACS). Methods We included 136 patients (71 men; 61.0 {+/-} 15.4 years) who underwent preoperative computed tomography (CT) and MRI within an 8-week period. Elliptical ROI HU values were measured at L2/3 and L4/5 for erector spinae, multifidus, and psoas major. MRI T2-hyperintense area fraction (Otsu thresholding) served as the fat infiltration reference. Linear mixed-effects (LME) models were used to assess the HU-T2 association and level-specific receiver operating characteristic (ROC) analyses (lower HU value side; n=136 per muscle-level) to identify thresholds for [≥]30% and [≥]50% infiltration criteria. Results Intraclass coefficients = 0.709 (HU) and 0.857 (T2 fraction); Goutallier weighted kappa = 0.579. In the overall LME, {beta} was -0.880 HU per 1% T2-fraction increase (95% confidence interval -0.935 to -0.825; marginal R2 =0.502); the association was steeper in multifidus ({beta} = -1.020) than in erector spinae ({beta} = -0.753). Psoas major (R = -0.226) was excluded from ROC analyses. Difference between L2/3 and L4/5 HU cutoffs was ~20 HU. The [≥]50% criterion revealed higher discrimination. Conclusions Elliptical ROI-based HU measurements may reliably screen paraspinal muscle fat infiltration in erector spinae and multifidus using standard PACS. Specific thresholds may allow practical preoperative evaluation without additional costs or radiation.
Courtens, J.; Muller, F. M.; Li, E. J.; Vanhove, C.; Vandenberghe, S.; Pantel, A. R.; Karp, J. S.; Daube-Witherspoon, M. E.
Show abstract
Dynamic positron emission tomography (PET) with long axial field-of-view (LAFOV) scanners enables multi-organ imaging and kinetic quantification beyond static (late-phase) imaging; however, the long times typically required for dynamic acquisitions remain clinically impractical. This study evaluates a deep learning (DL) framework to enable abbreviated dynamic PET acquisitions, comparing single-time-window (STW, early dynamic data only) and dual-time-window (DTW, early dynamic data plus a late 5-min static frame) protocols with early dynamic scan durations of 5-30 min and dose levels ranging from 360 MBq to 18 MBq. Seventeen 60-min dynamic [18F]FDG datasets were first motion-corrected using a staggered FALCON pipeline and then used to train and test a spatiotemporal DL model for autoregressive frame prediction. Performance was assessed across the full quantitative workflow, from DL-predicted frames and time-activity curves to organ-based kinetic modeling and voxel-wise parametric imaging in multiple tissues and two patient cohorts. DTW protocols consistently outperformed STW, better preserving late-phase kinetics. For a 15-min early dynamic scan, adding a late 5-min scan reduced mean absolute Ki difference from 23% (STW) to 17% (DTW) in the liver and from 26% to 15% in the thalamus. DTW + DL further reduced errors to [≤]10% in the liver, thalamus, and breast lesion, and 16% in muscle. Our recommended protocol, 15-min early dynamic scan plus a 5-min late scan with DL, remained robust to up to a 5-fold dose reduction (~74 MBq). Overall, these findings support DL-enabled abbreviated, low-dose dynamic LAFOV PET as a clinically feasible approach for accurate kinetic quantification
Bourne, R. M.; Arhatari, B.; Watson, G.; Gureyev, T.; Phipps, A.; Dowland, S.; Kurniawan, N.; Sved, P.
Show abstract
Formalin-fixed prostate tissue samples were imaged by propagation-based synchrotron phase contrast micro computed tomography ({micro}CT) with a 3D spatial resolution of ca. 3 {micro}m. Post-{micro}CT, samples were prepared for histology with sections close to coplanar with the transverse {micro}CT image planes. Haematoxylin and eosin stained sections were examined by an expert prostate histopathologist and compared qualitatively with corresponding {micro}CT-visible microstructure features. There is potential for {micro}CT to provide complimentary information to conventional histology and light microscopy without the need for preparation of stained thin sections. For the imaging conditions and spatial resolution of our study, {micro}CT may provide tissue architectural features similar to those used in Gleason grading, albeit without clear subcellular microstructure detail. At the spatial resolution of our study {micro}CT may provide novel 3D microstructure information for validation of diffusion weighted magnetic resonance imaging (MRI) methods. As an example, we demonstrate a qualitative correlation between {micro}CT-derived stromal fibre orientation and preferential water diffusion direction measured by diffusion tensor MRI microscopy of the same sample.
Mostafavi, S.; Shanbhag, A.; Ramirez, G.; Lemley, M.; Miller, R. J. H.; Chareonthaitawee, P.; Liang, J. X.; Dey, D.; Kavanagh, P. B.; Slipczuk, L.; Travin, M. I.; Alexanderson, E.; Carvajal Juarez, I.; Packard, R. R.; Al-Mallah, M. H.; Einstein, A. J.; Ruddy, T. D.; deKemp, R. A.; Boczar, K.; Feher, A.; Buechel, R. R.; Acampa, W.; Knight, S.; Le, V. T.; Rosamond, T. L.; Berman, D. S.; Di Carli, M. F.; Slomka, P.
Show abstract
Background: Positron emission tomography (PET) myocardial perfusion imaging (MPI) provides complementary information on perfusion, myocardial blood flow and ventricular function. While these markers are often considered collectively during interpretation, their quantitative integration with imaging and clinical data into a unified predictive framework remains limited. We developed a multimodal artificial intelligence framework that combines PET polar maps with quantitative imaging and clinical features to improve obstructive coronary artery disease (CAD) detection. Methods: We retrospectively analyzed the multicenter REFINE PET registry. Among 38,682 PET MPI studies from 14 sites, 2,833 patients without known prior CAD underwent invasive coronary angiography within 180 days. Obstructive CAD was defined as >=50% left main stenosis or >=70% stenosis in other major epicardial coronary arteries. We developed a two-stage contrastive learning framework to learn multimodal PET representations from studies without angiographic labels and transfer them to supervised CAD prediction. In Stage 1, PET image and tabular encoders were pretrained on 12,225 PET MPI studies from eight development sites using 15-channel PET polar maps, quantitative PET perfusion, flow and gated functional measures, and clinical variables. In Stage 2, the pretrained encoders and a lightweight classification head were fine-tuned in 968 angiography-labeled patients, using lower encoder learning rates to limit overfitting. The model was externally validated for angiographically defined obstructive CAD detection in 1,865 patients from six independent sites and compared with standard PET MPI metrics. Results: The prevalence of obstructive CAD was 60% in the training cohort (66% male, median age of 70 years [63, 77]), and 55% in the external validation cohort (64% male, median age of 67 years [60-74]). In external validation, the AI model achieved an AUC of 0.85 (95% confidence interval (CI), 0.83-0.87) for obstructive CAD detection and outperformed conventional quantitative PET metrics (all P < 0.001). At a specificity matched to visual summed stress score, the AI model achieved higher sensitivity (89% [95% CI, 87-91] versus 85% [95% CI, 82-87]) and negative predictive value (81% [95% CI, 77-84] versus 73% [95% CI, 69-77]; both p<0.001). The overall net reclassification improvement was 8.9% (95% CI, 4.2-13.6%; p = 0.001). Conclusions: Multimodal contrastive pretraining improved obstructive CAD detection from PET imaging beyond conventional perfusion-based scoring in independent multisite external validation.
Naidu, J.; Muralidharan, S.; Prashani, A.; Baskaradoss, V.
Show abstract
Objectives: To test whether radiology report evaluation metrics distinguish clinically meaningful errors from textual changes and align with radiologist-assessed error burden. Methods: Cross-dataset evaluation used ReXErr-v1 (2,708 report pairs; 5,724 paired error sentences) and 100 RadEvalX report pairs with expert error counts. BLEU-4, ROUGE-L and METEOR were assessed in ReXErr-v1; RadEvalX analyses included these plus BERTScore, CheXbert, RadGraph F1 and RadCliQ. Outcomes were ReXErr-v1 pairwise win rate and AUROC for clinical-content versus linguistic errors, and RadEvalX Spearman correlation with clinically significant error count and AUROC for any significant error. Confidence intervals used 10,000 clustered percentile bootstrap resamples; Holm adjustment-controlled multiplicity. Results: ReXErr-v1 paired-sentence win rates were 0.986 for BLEU-4, 0.999 for ROUGE-L and 0.998 for METEOR, but discrimination of clinical-content from linguistic errors was modest (AUROC 0.609-0.620). Penalty magnitude was strongly associated with textual change after adjustment for error type (normalised character edit distance coefficient 0.746; 95% CI 0.705-0.788; P<0.001). In RadEvalX, CheXbert showed the highest correlation with clinically significant errors (rho=0.413; 95% CI 0.223-0.578) and highest AUROC (0.742; 95% CI 0.638-0.836). Conclusions: Near-ceiling sensitivity to textual corruption did not imply sensitivity to clinical significance. CheXbert showed the highest alignment with expert error assessment, although pairwise superiority was not demonstrated over all comparators and performance remained moderate.
Bondarenko, M.; Qi, K.; Nowroozi, A.; Kim, J.; Kunzang, B.; Lee, A.; Liu, J.; Tran, N.; Weng, S.; Vella, M.; Chaudhari, G.; Schnizler, T.; Innanje, A.; Chen, T.; Sohn, J. H.
Show abstract
Background: Prediction of subsolid pulmonary nodule (SSN) progression from baseline CT may improve risk stratification and surveillance planning, but prior approaches have largely relied on fixed follow-up intervals. Methods: This retrospective single-center study evaluated interval-aware temporal imaging models for predicting future SSN growth and morphology across heterogeneous surveillance durations. A total of 24,946 longitudinal scan pairings derived from 2,543 clinician-reviewed SSNs in 426 patients were analyzed. A discriminative deep learning model predicted interval growth from baseline CT, segmentation masks, and interscan interval information, while a temporally conditioned generative model predicted future lesion morphology. Results: The discriminative model achieved an area under the receiver operating characteristic curve of 0.772 (95% confidence interval: 0.704-0.818), with sensitivity of 80.2% and specificity of 58.7% on the test cohort. The generative model predicted future lesion morphology with a Dice similarity coefficient of 0.706 +/-0.186. Prediction performance decreased with increasing follow-up duration, although both models generalized across intervals ranging from months to years. Conclusion: Interval-aware temporal imaging models enable the prediction of future SSN growth and morphology from baseline CT while accounting for variable surveillance intervals. These findings suggest a framework for time-aware, personalized risk assessment that may support individualized surveillance strategies and future AI-assisted management of pulmonary adenocarcinoma spectrum lesions.